Micron Document
<!DOCTYPE html>
<html class="client-nojs vector-feature-night-mode-disabled vector-feature-language-in-header-enabled vector-feature-language-in-main-page-header-disabled vector-feature-page-tools-pinned-disabled vector-feature-toc-pinned-clientpref-1 vector-feature-main-menu-pinned-disabled vector-feature-limited-width-clientpref-1 vector-feature-limited-width-content-enabled vector-feature-custom-font-size-clientpref-1 vector-feature-appearance-pinned-clientpref-1 vector-sticky-header-enabled" lang="en" dir="ltr"><head>
<meta charset="UTF-8">
<title>Speech processing</title>
<meta name="viewport" content="width=device-width, initial-scale=1.0">
<link rel="canonical" href="https://en.wikipedia.org/wiki/Speech_processing"> <link href="./mw/ext.cite.styles.css" rel="stylesheet" type="text/css">
<link href="./mw/ext.math.styles.css" rel="stylesheet" type="text/css">
<link href="./mw/skins.vector.icons.css" rel="stylesheet" type="text/css">
<link href="./mw/skins.vector.search.codex.styles.css" rel="stylesheet" type="text/css">
<link href="./mw/skins.vector.styles.css" rel="stylesheet" type="text/css">
<link href="./mw/user.styles.css" rel="stylesheet" type="text/css">
<meta name="ResourceLoaderDynamicStyles" content="">
<link rel="stylesheet" type="text/css" href="./mw/site.styles.css">
<link rel="stylesheet" type="text/css" href="./mw/noscript.css">
<link rel="stylesheet" type="text/css" href="./footer.css">
<link rel="stylesheet" type="text/css" href="./vector-2022.css">
</head>
<body class="skin--responsive skin-vector skin-vector-search-vue mediawiki ltr sitedir-ltr mw-hide-empty-elt ns-0 ns-subject page-Speech_processing rootpage-Speech_processing skin-vector-2022 action-view">
<div class="mw-page-container">
<div class="mw-page-container-inner">
<div class="mw-content-container">
<main id="content" class="mw-body">
<header class="mw-body-header vector-page-titlebar">
<h1 id="firstHeading" class="firstHeading mw-first-heading">
<span id="openzim-page-title" class="mw-page-title-main"><span class="mw-page-title-main">Speech processing</span></span>
</h1>
</header>
<a id="top"></a>
<div id="bodyContent" class="vector-body ve-init-mw-desktopArticleTarget-targetContainer" aria-labelledby="firstHeading" data-mw-ve-target-container="">
<div id="mw-content-text" class="mw-body-content mw-content-ltr" lang="en" dir="ltr"><div class="mw-content-ltr mw-parser-output" lang="en" dir="ltr">
<style data-mw-deduplicate="TemplateStyles:r1236090951">
/* start https://en.wikipedia.org/ */


.mw-parser-output .hatnote{font-style:italic}.mw-parser-output div.hatnote{padding-left:1.6em;margin-bottom:0.5em}.mw-parser-output .hatnote i{font-style:normal}.mw-parser-output .hatnote+link+.hatnote{margin-top:-0.5em}@media print{body.ns-0 .mw-parser-output .hatnote{display:none!important}}


/* end https://en.wikipedia.org/ */
</style><div role="note" class="hatnote navigation-not-searchable">This article is about electronic speech processing. For speech processing in the human brain, see <a href="Language_processing_in_the_brain" title="Language processing in the brain">Language processing in the brain</a>.</div>
<p><b>Speech processing</b> is the study of <a href="Speech_communication" class="mw-redirect" title="Speech communication">speech</a> <a href="Signal_(information_theory)" class="mw-redirect" title="Signal (information theory)">signals</a> and the processing methods of signals. The signals are usually processed in a <a href="Digital_data" title="Digital data">digital</a> representation, so speech processing can be regarded as a special case of <a href="Digital_signal_processing" title="Digital signal processing">digital signal processing</a>, applied to <a href="Audio_signal" title="Audio signal">speech signals</a>. Aspects of speech processing includes the acquisition, manipulation, storage, transfer and output of speech signals. Different speech processing tasks include <a href="Speech_recognition" title="Speech recognition">speech recognition</a>, <a href="Speech_synthesis" title="Speech synthesis">speech synthesis</a>, <a href="Speaker_diarization" class="mw-redirect" title="Speaker diarization">speaker diarization</a>, <a href="Speech_enhancement" title="Speech enhancement">speech enhancement</a>, <a href="Speaker_recognition" title="Speaker recognition">speaker recognition</a>, etc.<sup id="cite_ref-1" class="reference"><a href="#cite_note-1"><span class="cite-bracket">[</span>1<span class="cite-bracket">]</span></a></sup>
</p>
<meta property="mw:PageProp/toc">
<div class="mw-heading mw-heading2"><h2 id="History">History</h2></div>
<p>Early attempts at speech processing and recognition were primarily focused on understanding a handful of simple <a href="Phonetics" title="Phonetics">phonetic</a> elements such as vowels. In 1952, three researchers at Bell Labs, Stephen. Balashek, R. Biddulph, and K. H. Davis, developed a system that could recognize digits spoken by a single speaker.<sup id="cite_ref-2" class="reference"><a href="#cite_note-2"><span class="cite-bracket">[</span>2<span class="cite-bracket">]</span></a></sup> Pioneering works in field of speech recognition using analysis of its spectrum were reported in the 1940s.<sup id="cite_ref-3" class="reference"><a href="#cite_note-3"><span class="cite-bracket">[</span>3<span class="cite-bracket">]</span></a></sup>
</p><p><a href="Linear_predictive_coding" title="Linear predictive coding">Linear predictive coding</a> (LPC), a speech processing algorithm, was first proposed by <a href="Fumitada_Itakura" title="Fumitada Itakura">Fumitada Itakura</a> of <a href="Nagoya_University" title="Nagoya University">Nagoya University</a> and Shuzo Saito of <a href="Nippon_Telegraph_and_Telephone" title="Nippon Telegraph and Telephone">Nippon Telegraph and Telephone</a> (NTT) in 1966.<sup id="cite_ref-Gray_4-0" class="reference"><a href="#cite_note-Gray-4"><span class="cite-bracket">[</span>4<span class="cite-bracket">]</span></a></sup> Further developments in LPC technology were made by <a href="Bishnu_S._Atal" title="Bishnu S. Atal">Bishnu S. Atal</a> and <a href="Manfred_R._Schroeder" title="Manfred R. Schroeder">Manfred R. Schroeder</a> at <a href="Bell_Labs" title="Bell Labs">Bell Labs</a> during the 1970s.<sup id="cite_ref-Gray_4-1" class="reference"><a href="#cite_note-Gray-4"><span class="cite-bracket">[</span>4<span class="cite-bracket">]</span></a></sup> LPC was the basis for <a href="Voice-over-IP" class="mw-redirect" title="Voice-over-IP">voice-over-IP</a> (VoIP) technology,<sup id="cite_ref-Gray_4-2" class="reference"><a href="#cite_note-Gray-4"><span class="cite-bracket">[</span>4<span class="cite-bracket">]</span></a></sup> as well as <a href="Speech_synthesizer" class="mw-redirect" title="Speech synthesizer">speech synthesizer</a> chips, such as the <a href="Texas_Instruments_LPC_Speech_Chips" title="Texas Instruments LPC Speech Chips">Texas Instruments LPC Speech Chips</a> used in the <a href="Speak_%26_Spell_(toy)" title="Speak &amp; Spell (toy)">Speak &amp; Spell</a> toys from 1978.<sup id="cite_ref-vintagecomputing_article_5-0" class="reference"><a href="#cite_note-vintagecomputing_article-5"><span class="cite-bracket">[</span>5<span class="cite-bracket">]</span></a></sup>
</p><p>One of the first commercially available speech recognition products was Dragon Dictate, released in 1990. In 1992, technology developed by <a href="Lawrence_Rabiner" title="Lawrence Rabiner">Lawrence Rabiner</a> and others at Bell Labs was used by <a href="AT%26T" title="AT&amp;T">AT&amp;T</a> in their Voice Recognition Call Processing service to route calls without a human operator. By this point, the vocabulary of these systems was larger than the average human vocabulary.<sup id="cite_ref-6" class="reference"><a href="#cite_note-6"><span class="cite-bracket">[</span>6<span class="cite-bracket">]</span></a></sup>
</p><p>By the early 2000s, the dominant speech processing strategy started to shift away from <a href="Hidden_Markov_model" title="Hidden Markov model">Hidden Markov Models</a> towards more modern <a href="Artificial_neural_network" class="mw-redirect" title="Artificial neural network">neural networks</a> and <a href="Deep_learning" title="Deep learning">deep learning</a>.<sup id="cite_ref-7" class="reference"><a href="#cite_note-7"><span class="cite-bracket">[</span>7<span class="cite-bracket">]</span></a></sup>
</p><p>In 2012, <a href="Geoffrey_Hinton" title="Geoffrey Hinton">Geoffrey Hinton</a> and his team at the <a href="University_of_Toronto" title="University of Toronto">University of Toronto</a> demonstrated that deep neural networks could significantly outperform traditional HMM-based systems on large vocabulary continuous speech recognition tasks. This breakthrough led to widespread adoption of deep learning techniques in the industry.<sup id="cite_ref-:0_8-0" class="reference"><a href="#cite_note-:0-8"><span class="cite-bracket">[</span>8<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-9" class="reference"><a href="#cite_note-9"><span class="cite-bracket">[</span>9<span class="cite-bracket">]</span></a></sup>
</p><p>By the mid-2010s, companies like <a href="Google" title="Google">Google</a>, <a href="Microsoft" title="Microsoft">Microsoft</a>, <a href="Amazon_(company)" title="Amazon (company)">Amazon</a>, and <a href="Apple_Inc." title="Apple Inc.">Apple</a> had integrated advanced speech recognition systems into their virtual assistants such as <a href="Google_Assistant" title="Google Assistant">Google Assistant</a>, <a href="Cortana_(virtual_assistant)" title="Cortana (virtual assistant)">Cortana</a>, <a href="Amazon_Alexa" title="Amazon Alexa">Alexa</a>, and <a href="Siri" title="Siri">Siri</a>.<sup id="cite_ref-10" class="reference"><a href="#cite_note-10"><span class="cite-bracket">[</span>10<span class="cite-bracket">]</span></a></sup> These systems utilized deep learning models to provide more natural and accurate voice interactions.
</p><p>The development of Transformer-based models, like Google's BERT (Bidirectional Encoder Representations from Transformers) and OpenAI's GPT (Generative Pre-trained Transformer), further pushed the boundaries of natural language processing and speech recognition. These models enabled more context-aware and semantically rich understanding of speech.<sup id="cite_ref-11" class="reference"><a href="#cite_note-11"><span class="cite-bracket">[</span>11<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-:0_8-1" class="reference"><a href="#cite_note-:0-8"><span class="cite-bracket">[</span>8<span class="cite-bracket">]</span></a></sup> In recent years, end-to-end speech recognition models have gained popularity. These models simplify the speech recognition pipeline by directly converting audio input into text output, bypassing intermediate steps like feature extraction and acoustic modeling. This approach has streamlined the development process and improved performance.<sup id="cite_ref-12" class="reference"><a href="#cite_note-12"><span class="cite-bracket">[</span>12<span class="cite-bracket">]</span></a></sup>
</p>
<div class="mw-heading mw-heading2"><h2 id="Techniques">Techniques</h2></div>
<div class="mw-heading mw-heading3"><h3 id="Dynamic_time_warping">Dynamic time warping</h3></div>
<div role="note" class="hatnote navigation-not-searchable">Main article: <a href="Dynamic_time_warping" title="Dynamic time warping">Dynamic time warping</a></div><p>Dynamic time warping (DTW) is an <a href="Algorithm" title="Algorithm">algorithm</a> for measuring similarity between two <a href="Time_series" title="Time series">temporal sequences</a>, which may vary in speed. In general, DTW is a method that calculates an <a href="Optimal_matching" title="Optimal matching">optimal match</a> between two given sequences (e.g. time series) with certain restriction and rules. The optimal match is denoted by the match that satisfies all the restrictions and the rules and that has the minimal cost, where the cost is computed as the sum of absolute differences, for each matched pair of indices, between their values.
</p><div class="mw-heading mw-heading3"><h3 id="Hidden_Markov_models">Hidden Markov models</h3></div>
<div role="note" class="hatnote navigation-not-searchable">Main article: <a href="Hidden_Markov_model" title="Hidden Markov model">Hidden Markov model</a></div><p>A hidden Markov model can be represented as the simplest <a href="Dynamic_Bayesian_network" title="Dynamic Bayesian network">dynamic Bayesian network</a>. The goal of the algorithm is to estimate a hidden variable x(t) given a list of observations y(t). By applying the <a href="Markov_property" title="Markov property">Markov property</a>, the <a href="Conditional_probability_distribution" title="Conditional probability distribution">conditional probability distribution</a> of the hidden variable <i>x</i>(<i>t</i>) at time <i>t</i>, given the values of the hidden variable <i>x</i> at all times, depends <i>only</i> on the value of the hidden variable <i>x</i>(<i>t</i> − 1). Similarly, the value of the observed variable <i>y</i>(<i>t</i>) only depends on the value of the hidden variable <i>x</i>(<i>t</i>) (both at time <i>t</i>).
</p><div class="mw-heading mw-heading3"><h3 id="Artificial_neural_networks">Artificial neural networks</h3></div>
<div role="note" class="hatnote navigation-not-searchable">Main article: <a href="Artificial_neural_network" class="mw-redirect" title="Artificial neural network">Artificial neural network</a></div><p>An artificial neural network (ANN) is based on a collection of connected units or nodes called <a href="Artificial_neuron" title="Artificial neuron">artificial neurons</a>, which loosely model the <a href="Neuron" title="Neuron">neurons</a> in a biological <a href="Brain" title="Brain">brain</a>. Each connection, like the <a href="Synapse" title="Synapse">synapses</a> in a biological <a href="Brain" title="Brain">brain</a>, can transmit a signal from one artificial neuron to another. An artificial neuron that receives a signal can process it and then signal additional artificial neurons connected to it. In common ANN implementations, the signal at a connection between artificial neurons is a <a href="Real_number" title="Real number">real number</a>, and the output of each artificial neuron is computed by some non-linear function of the sum of its inputs.
</p><div class="mw-heading mw-heading3"><h3 id="Phase-aware_processing">Phase-aware processing</h3></div>
<p>Phase is often assumed to be random, but contains useful information. Wrapping of phase:<sup id="cite_ref-limits_13-0" class="reference"><a href="#cite_note-limits-13"><span class="cite-bracket">[</span>13<span class="cite-bracket">]</span></a></sup> can be introduced due to periodical jumps on <span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle 2\pi }">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<mn>2</mn>
<mi>π<!-- π --></mi>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle 2\pi }</annotation>
</semantics>
</math></span><img src="./73efd1f6493490b058097060a572606d2c550a06.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.338ex; width:2.494ex; height:2.176ex;" alt="{\displaystyle 2\pi }" loading="lazy"></span>. Phase unwrapping (see,<sup id="cite_ref-14" class="reference"><a href="#cite_note-14"><span class="cite-bracket">[</span>14<span class="cite-bracket">]</span></a></sup> Chapter 2.3; <a href="Instantaneous_phase_and_frequency" title="Instantaneous phase and frequency">Instantaneous phase and frequency</a>), it can be expressed as:<sup id="cite_ref-limits_13-1" class="reference"><a href="#cite_note-limits-13"><span class="cite-bracket">[</span>13<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-vonMises_15-0" class="reference"><a href="#cite_note-vonMises-15"><span class="cite-bracket">[</span>15<span class="cite-bracket">]</span></a></sup>
<span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle \phi (h,l)=\phi _{lin}(h,l)+\Psi (h,l)}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<mi>ϕ<!-- ϕ --></mi>
<mo stretchy="false">(</mo>
<mi>h</mi>
<mo>,</mo>
<mi>l</mi>
<mo stretchy="false">)</mo>
<mo>=</mo>
<msub>
<mi>ϕ<!-- ϕ --></mi>
<mrow class="MJX-TeXAtom-ORD">
<mi>l</mi>
<mi>i</mi>
<mi>n</mi>
</mrow>
</msub>
<mo stretchy="false">(</mo>
<mi>h</mi>
<mo>,</mo>
<mi>l</mi>
<mo stretchy="false">)</mo>
<mo>+</mo>
<mi mathvariant="normal">Ψ<!-- Ψ --></mi>
<mo stretchy="false">(</mo>
<mi>h</mi>
<mo>,</mo>
<mi>l</mi>
<mo stretchy="false">)</mo>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle \phi (h,l)=\phi _{lin}(h,l)+\Psi (h,l)}</annotation>
</semantics>
</math></span><img src="./27b9b179b1f238a7ccfc6078665851084b3b5201.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.838ex; width:27.42ex; height:2.843ex;" alt="{\displaystyle \phi (h,l)=\phi _{lin}(h,l)+\Psi (h,l)}" loading="lazy"></span>, where <span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle \phi _{lin}(h,l)=\omega _{0}(l'){}_{\Delta }t}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<msub>
<mi>ϕ<!-- ϕ --></mi>
<mrow class="MJX-TeXAtom-ORD">
<mi>l</mi>
<mi>i</mi>
<mi>n</mi>
</mrow>
</msub>
<mo stretchy="false">(</mo>
<mi>h</mi>
<mo>,</mo>
<mi>l</mi>
<mo stretchy="false">)</mo>
<mo>=</mo>
<msub>
<mi>ω<!-- ω --></mi>
<mrow class="MJX-TeXAtom-ORD">
<mn>0</mn>
</mrow>
</msub>
<mo stretchy="false">(</mo>
<msup>
<mi>l</mi>
<mo>′</mo>
</msup>
<mo stretchy="false">)</mo>
<msub>
<mrow class="MJX-TeXAtom-ORD">

</mrow>
<mrow class="MJX-TeXAtom-ORD">
<mi mathvariant="normal">Δ<!-- Δ --></mi>
</mrow>
</msub>
<mi>t</mi>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle \phi _{lin}(h,l)=\omega _{0}(l'){}_{\Delta }t}</annotation>
</semantics>
</math></span><img src="./1d2c51a02117724ca424de1acd659f1702f9c3d4.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.838ex; width:19.764ex; height:3.009ex;" alt="{\displaystyle \phi _{lin}(h,l)=\omega _{0}(l'){}_{\Delta }t}" loading="lazy"></span> is linear phase (<span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle {}_{\Delta }t}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<msub>
<mrow class="MJX-TeXAtom-ORD">

</mrow>
<mrow class="MJX-TeXAtom-ORD">
<mi mathvariant="normal">Δ<!-- Δ --></mi>
</mrow>
</msub>
<mi>t</mi>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle {}_{\Delta }t}</annotation>
</semantics>
</math></span><img src="./88ff57f69723ccb8153df6cdb30b478fed75e41d.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.671ex; width:2.441ex; height:2.343ex;" alt="{\displaystyle {}_{\Delta }t}" loading="lazy"></span> is temporal shift at each frame of analysis), <span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle \Psi (h,l)}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<mi mathvariant="normal">Ψ<!-- Ψ --></mi>
<mo stretchy="false">(</mo>
<mi>h</mi>
<mo>,</mo>
<mi>l</mi>
<mo stretchy="false">)</mo>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle \Psi (h,l)}</annotation>
</semantics>
</math></span><img src="./38ddf67a07c5d6a09c7131a612edeeb38ac022e3.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.838ex; width:6.684ex; height:2.843ex;" alt="{\displaystyle \Psi (h,l)}" loading="lazy"></span> is phase contribution of the vocal tract and phase source.<sup id="cite_ref-vonMises_15-1" class="reference"><a href="#cite_note-vonMises-15"><span class="cite-bracket">[</span>15<span class="cite-bracket">]</span></a></sup>
Obtained phase estimations can be used for noise reduction: temporal smoothing of instantaneous phase <sup id="cite_ref-16" class="reference"><a href="#cite_note-16"><span class="cite-bracket">[</span>16<span class="cite-bracket">]</span></a></sup> and its derivatives by time (<a href="Instantaneous_phase_and_frequency" title="Instantaneous phase and frequency">instantaneous frequency</a>) and frequency (<a href="Group_delay_and_phase_delay" title="Group delay and phase delay">group delay</a>),<sup id="cite_ref-Advances_17-0" class="reference"><a href="#cite_note-Advances-17"><span class="cite-bracket">[</span>17<span class="cite-bracket">]</span></a></sup> smoothing of phase across frequency.<sup id="cite_ref-Advances_17-1" class="reference"><a href="#cite_note-Advances-17"><span class="cite-bracket">[</span>17<span class="cite-bracket">]</span></a></sup> Joined amplitude and phase estimators can recover speech more accurately basing on assumption of von Mises distribution of phase.<sup id="cite_ref-vonMises_15-2" class="reference"><a href="#cite_note-vonMises-15"><span class="cite-bracket">[</span>15<span class="cite-bracket">]</span></a></sup>
</p>
<div class="mw-heading mw-heading2"><h2 id="Applications">Applications</h2></div>
<ul><li><a href="Interactive_voice_response" title="Interactive voice response">Interactive voice response</a></li>
<li><a href="Virtual_assistant" title="Virtual assistant">Virtual Assistants</a></li>
<li><a href="Speaker_recognition" title="Speaker recognition">Voice Identification</a></li>
<li><a href="Emotion_recognition" title="Emotion recognition">Emotion Recognition</a></li>
<li>Call Center Automation</li>
<li><a href="Robotics" title="Robotics">Robotics</a></li></ul>
<div class="mw-heading mw-heading2"><h2 id="See_also">See also</h2></div>
<ul><li><a href="Computational_audiology" title="Computational audiology">Computational audiology</a></li>
<li><a href="Neurocomputational_speech_processing" title="Neurocomputational speech processing">Neurocomputational speech processing</a></li>
<li><a href="Speech_coding" title="Speech coding">Speech coding</a></li>
<li><a href="Speech_technology" title="Speech technology">Speech technology</a></li>
<li><a href="Natural_language_processing" title="Natural language processing">Natural Language Processing</a></li></ul>
<div class="mw-heading mw-heading2"><h2 id="References">References</h2></div>
<style data-mw-deduplicate="TemplateStyles:r1239543626">
/* start https://en.wikipedia.org/ */


.mw-parser-output .reflist{margin-bottom:0.5em;list-style-type:decimal}@media screen{.mw-parser-output .reflist{font-size:90%}}.mw-parser-output .reflist .references{font-size:100%;margin-bottom:0;list-style-type:inherit}.mw-parser-output .reflist-columns-2{column-width:30em}.mw-parser-output .reflist-columns-3{column-width:25em}.mw-parser-output .reflist-columns{margin-top:0.3em}.mw-parser-output .reflist-columns ol{margin-top:0}.mw-parser-output .reflist-columns li{page-break-inside:avoid;break-inside:avoid-column}.mw-parser-output .reflist-upper-alpha{list-style-type:upper-alpha}.mw-parser-output .reflist-upper-roman{list-style-type:upper-roman}.mw-parser-output .reflist-lower-alpha{list-style-type:lower-alpha}.mw-parser-output .reflist-lower-greek{list-style-type:lower-greek}.mw-parser-output .reflist-lower-roman{list-style-type:lower-roman}


/* end https://en.wikipedia.org/ */
</style><div class="reflist">
<div class="mw-references-wrap mw-references-columns"><ol class="references">
<li id="cite_note-1"><span class="mw-cite-backlink"><b><a href="#cite_ref-1">^</a></b></span> <span class="reference-text"><style data-mw-deduplicate="TemplateStyles:r1238218222">
/* start https://en.wikipedia.org/ */


.mw-parser-output cite.citation{font-style:inherit;word-wrap:break-word}.mw-parser-output .citation q{quotes:"\"""\"""'""'"}.mw-parser-output .citation:target{background-color:rgba(0,127,255,0.133)}.mw-parser-output .id-lock-free.id-lock-free a{background:url("./mw/Lock-green.svg")right 0.1em center/9px no-repeat}.mw-parser-output .id-lock-limited.id-lock-limited a,.mw-parser-output .id-lock-registration.id-lock-registration a{background:url("./mw/Lock-gray-alt-2.svg")right 0.1em center/9px no-repeat}.mw-parser-output .id-lock-subscription.id-lock-subscription a{background:url("./mw/Lock-red-alt-2.svg")right 0.1em center/9px no-repeat}.mw-parser-output .cs1-ws-icon a{background:url("./mw/Wikisource-logo.svg")right 0.1em center/12px no-repeat}body:not(.skin-timeless):not(.skin-minerva) .mw-parser-output .id-lock-free a,body:not(.skin-timeless):not(.skin-minerva) .mw-parser-output .id-lock-limited a,body:not(.skin-timeless):not(.skin-minerva) .mw-parser-output .id-lock-registration a,body:not(.skin-timeless):not(.skin-minerva) .mw-parser-output .id-lock-subscription a,body:not(.skin-timeless):not(.skin-minerva) .mw-parser-output .cs1-ws-icon a{background-size:contain;padding:0 1em 0 0}.mw-parser-output .cs1-code{color:inherit;background:inherit;border:none;padding:inherit}.mw-parser-output .cs1-hidden-error{display:none;color:var(--color-error,#d33)}.mw-parser-output .cs1-visible-error{color:var(--color-error,#d33)}.mw-parser-output .cs1-maint{display:none;color:#085;margin-left:0.3em}.mw-parser-output .cs1-kern-left{padding-left:0.2em}.mw-parser-output .cs1-kern-right{padding-right:0.2em}.mw-parser-output .citation .mw-selflink{font-weight:inherit}@media screen{.mw-parser-output .cs1-format{font-size:95%}html.skin-theme-clientpref-night .mw-parser-output .cs1-maint{color:#18911f}}@media screen and (prefers-color-scheme:dark){html.skin-theme-clientpref-os .mw-parser-output .cs1-maint{color:#18911f}}


/* end https://en.wikipedia.org/ */
</style><cite id="CITEREFSahidullahPatinoCornellYin2019" class="citation arxiv cs1">Sahidullah, Md; Patino, Jose; Cornell, Samuele; Yin, Ruiking; Sivasankaran, Sunit; Bredin, Herve; Korshunov, Pavel; Brutti, Alessio; Serizel, Romain; Vincent, Emmanuel; Evans, Nicholas; Marcel, Sebastien; Squartini, Stefano; Barras, Claude (2019-11-06). "The Speed Submission to DIHARD II: Contributions &amp; Lessons Learned". <a href="ArXiv_(identifier)" class="mw-redirect" title="ArXiv (identifier)">arXiv</a>:<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://arxiv.org/abs/1911.02388">1911.02388</a></span> [<a rel="nofollow" class="external text" href="https://arxiv.org/archive/eess.AS">eess.AS</a>].</cite></span>
</li>
<li id="cite_note-2"><span class="mw-cite-backlink"><b><a href="#cite_ref-2">^</a></b></span> <span class="reference-text"><cite id="CITEREFJuangRabiner2006" class="citation cs2">Juang, B.-H.; Rabiner, L.R. (2006), "Speech Recognition, Automatic: History", <i>Encyclopedia of Language &amp; Linguistics</i>, Elsevier, pp.&nbsp;<span class="nowrap">806–</span>819, <a href="Doi_(identifier)" class="mw-redirect" title="Doi (identifier)">doi</a>:<a rel="nofollow" class="external text" href="https://doi.org/10.1016%2Fb0-08-044854-2%2F00906-8">10.1016/b0-08-044854-2/00906-8</a>, <a href="ISBN_(identifier)" class="mw-redirect" title="ISBN (identifier)">ISBN</a>&nbsp;<bdi>9780080448541</bdi></cite></span>
</li>
<li id="cite_note-3"><span class="mw-cite-backlink"><b><a href="#cite_ref-3">^</a></b></span> <span class="reference-text"><cite id="CITEREFMyasnikovMyasnikova1970" class="citation book cs1 cs1-prop-foreign-lang-source">Myasnikov, L. L.; Myasnikova, Ye. N. (1970). <i>Automatic recognition of sound pattern</i> (in Russian). Leningrad: Energiya.</cite></span>
</li>
<li id="cite_note-Gray-4"><span class="mw-cite-backlink">^ <a href="#cite_ref-Gray_4-0"><sup><i><b>a</b></i></sup></a> <a href="#cite_ref-Gray_4-1"><sup><i><b>b</b></i></sup></a> <a href="#cite_ref-Gray_4-2"><sup><i><b>c</b></i></sup></a></span> <span class="reference-text"><cite id="CITEREFGray2010" class="citation journal cs1">Gray, Robert M. (2010). <a rel="nofollow" class="external text" href="https://ee.stanford.edu/~gray/lpcip.pdf">"A History of Realtime Digital Speech on Packet Networks: Part II of Linear Predictive Coding and the Internet Protocol"</a> <span class="cs1-format">(PDF)</span>. <i>Found. Trends Signal Process</i>. <b>3</b> (4): <span class="nowrap">203–</span>303. <a href="Doi_(identifier)" class="mw-redirect" title="Doi (identifier)">doi</a>:<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://doi.org/10.1561%2F2000000036">10.1561/2000000036</a></span>. <a href="ISSN_(identifier)" class="mw-redirect" title="ISSN (identifier)">ISSN</a>&nbsp;<a rel="nofollow" class="external text" href="https://search.worldcat.org/issn/1932-8346">1932-8346</a>.</cite></span>
</li>
<li id="cite_note-vintagecomputing_article-5"><span class="mw-cite-backlink"><b><a href="#cite_ref-vintagecomputing_article_5-0">^</a></b></span> <span class="reference-text"><cite class="citation web cs1"><a rel="nofollow" class="external text" href="http://www.vintagecomputing.com/index.php/archives/528">"VC&amp;G - VC&amp;G Interview: 30 Years Later, Richard Wiggins Talks Speak &amp; Spell Development"</a>.</cite></span>
</li>
<li id="cite_note-6"><span class="mw-cite-backlink"><b><a href="#cite_ref-6">^</a></b></span> <span class="reference-text"><cite id="CITEREFHuangBakerReddy2014" class="citation journal cs1">Huang, Xuedong; Baker, James; Reddy, Raj (2014-01-01). "A historical perspective of speech recognition". <i>Communications of the ACM</i>. <b>57</b> (1): <span class="nowrap">94–</span>103. <a href="Doi_(identifier)" class="mw-redirect" title="Doi (identifier)">doi</a>:<a rel="nofollow" class="external text" href="https://doi.org/10.1145%2F2500887">10.1145/2500887</a>. <a href="ISSN_(identifier)" class="mw-redirect" title="ISSN (identifier)">ISSN</a>&nbsp;<a rel="nofollow" class="external text" href="https://search.worldcat.org/issn/0001-0782">0001-0782</a>. <a href="S2CID_(identifier)" class="mw-redirect" title="S2CID (identifier)">S2CID</a>&nbsp;<a rel="nofollow" class="external text" href="https://api.semanticscholar.org/CorpusID:6175701">6175701</a>.</cite></span>
</li>
<li id="cite_note-7"><span class="mw-cite-backlink"><b><a href="#cite_ref-7">^</a></b></span> <span class="reference-text"><cite id="CITEREFFurui2005" class="citation journal cs1">Furui, Sadaoki (2005). <a rel="nofollow" class="external text" href="https://doi.org/10.37936%2Fecti-cit.200512.51834">"50 Years of Progress in Speech and Speaker Recognition Research"</a>. <i>ECTI Transactions on Computer and Information Technology</i>. <b>1</b> (2): <span class="nowrap">64–</span>74. <a href="Doi_(identifier)" class="mw-redirect" title="Doi (identifier)">doi</a>:<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://doi.org/10.37936%2Fecti-cit.200512.51834">10.37936/ecti-cit.200512.51834</a></span>. <a href="ISSN_(identifier)" class="mw-redirect" title="ISSN (identifier)">ISSN</a>&nbsp;<a rel="nofollow" class="external text" href="https://search.worldcat.org/issn/2286-9131">2286-9131</a>.</cite></span>
</li>
<li id="cite_note-:0-8"><span class="mw-cite-backlink">^ <a href="#cite_ref-:0_8-0"><sup><i><b>a</b></i></sup></a> <a href="#cite_ref-:0_8-1"><sup><i><b>b</b></i></sup></a></span> <span class="reference-text"><cite class="citation news cs1"><a rel="nofollow" class="external text" href="https://www.cs.toronto.edu/~hinton/absps/DNN-2012-proof.pdf?form=MG0AV3">"Deep Neural Networks for Acoustic Modeling in Speech Recognition"</a> <span class="cs1-format">(PDF)</span>. 2019-07-23<span class="reference-accessdate">. Retrieved <span class="nowrap">2024-11-05</span></span>.</cite></span>
</li>
<li id="cite_note-9"><span class="mw-cite-backlink"><b><a href="#cite_ref-9">^</a></b></span> <span class="reference-text"><cite class="citation news cs1"><a rel="nofollow" class="external text" href="https://www.cs.toronto.edu/~hinton/absps/DRNN_speech.pdf?form=MG0AV3">"SPEECH RECOGNITION WITH DEEP RECURRENT NEURAL NETWORKS"</a> <span class="cs1-format">(PDF)</span>. 2019-07-23<span class="reference-accessdate">. Retrieved <span class="nowrap">2024-11-05</span></span>.</cite></span>
</li>
<li id="cite_note-10"><span class="mw-cite-backlink"><b><a href="#cite_ref-10">^</a></b></span> <span class="reference-text"><cite id="CITEREFHoy2018" class="citation journal cs1">Hoy, Matthew B. (2018). "Alexa, Siri, Cortana, and More: An Introduction to Voice Assistants". <i>Medical Reference Services Quarterly</i>. <b>37</b> (1): <span class="nowrap">81–</span>88. <a href="Doi_(identifier)" class="mw-redirect" title="Doi (identifier)">doi</a>:<a rel="nofollow" class="external text" href="https://doi.org/10.1080%2F02763869.2018.1404391">10.1080/02763869.2018.1404391</a>. <a href="ISSN_(identifier)" class="mw-redirect" title="ISSN (identifier)">ISSN</a>&nbsp;<a rel="nofollow" class="external text" href="https://search.worldcat.org/issn/1540-9597">1540-9597</a>. <a href="PMID_(identifier)" class="mw-redirect" title="PMID (identifier)">PMID</a>&nbsp;<a rel="nofollow" class="external text" href="https://pubmed.ncbi.nlm.nih.gov/29327988">29327988</a>.</cite></span>
</li>
<li id="cite_note-11"><span class="mw-cite-backlink"><b><a href="#cite_ref-11">^</a></b></span> <span class="reference-text"><cite class="citation web cs1 cs1-prop-foreign-lang-source"><a rel="nofollow" class="external text" href="https://vbee.vn">"Vbee"</a>. <i>vbee.vn</i> (in Vietnamese)<span class="reference-accessdate">. Retrieved <span class="nowrap">2024-11-05</span></span>.</cite></span>
</li>
<li id="cite_note-12"><span class="mw-cite-backlink"><b><a href="#cite_ref-12">^</a></b></span> <span class="reference-text"><cite id="CITEREFHagiwara2021" class="citation book cs1">Hagiwara, Masato (2021-12-21). <a rel="nofollow" class="external text" href="https://books.google.com/books?id=Ye9MEAAAQBAJ"><i>Real-World Natural Language Processing: Practical applications with deep learning</i></a>. Simon and Schuster. <a href="ISBN_(identifier)" class="mw-redirect" title="ISBN (identifier)">ISBN</a>&nbsp;<bdi>978-1-63835-039-2</bdi>.</cite></span>
</li>
<li id="cite_note-limits-13"><span class="mw-cite-backlink">^ <a href="#cite_ref-limits_13-0"><sup><i><b>a</b></i></sup></a> <a href="#cite_ref-limits_13-1"><sup><i><b>b</b></i></sup></a></span> <span class="reference-text"><cite id="CITEREFMowlaeeKulmer2015" class="citation journal cs1">Mowlaee, Pejman; Kulmer, Josef (August 2015). "Phase Estimation in Single-Channel Speech Enhancement: Limits-Potential". <i>IEEE/ACM Transactions on Audio, Speech, and Language Processing</i>. <b>23</b> (8): <span class="nowrap">1283–</span>1294. <a href="Doi_(identifier)" class="mw-redirect" title="Doi (identifier)">doi</a>:<a rel="nofollow" class="external text" href="https://doi.org/10.1109%2FTASLP.2015.2430820">10.1109/TASLP.2015.2430820</a>. <a href="ISSN_(identifier)" class="mw-redirect" title="ISSN (identifier)">ISSN</a>&nbsp;<a rel="nofollow" class="external text" href="https://search.worldcat.org/issn/2329-9290">2329-9290</a>. <a href="S2CID_(identifier)" class="mw-redirect" title="S2CID (identifier)">S2CID</a>&nbsp;<a rel="nofollow" class="external text" href="https://api.semanticscholar.org/CorpusID:13058142">13058142</a>.</cite></span>
</li>
<li id="cite_note-14"><span class="mw-cite-backlink"><b><a href="#cite_ref-14">^</a></b></span> <span class="reference-text"><cite id="CITEREFMowlaeeKulmerStahlMayer2017" class="citation book cs1">Mowlaee, Pejman; Kulmer, Josef; Stahl, Johannes; Mayer, Florian (2017). <i>Single channel phase-aware signal processing in speech communication: theory and practice</i>. Chichester: Wiley. <a href="ISBN_(identifier)" class="mw-redirect" title="ISBN (identifier)">ISBN</a>&nbsp;<bdi>978-1-119-23882-9</bdi>.</cite></span>
</li>
<li id="cite_note-vonMises-15"><span class="mw-cite-backlink">^ <a href="#cite_ref-vonMises_15-0"><sup><i><b>a</b></i></sup></a> <a href="#cite_ref-vonMises_15-1"><sup><i><b>b</b></i></sup></a> <a href="#cite_ref-vonMises_15-2"><sup><i><b>c</b></i></sup></a></span> <span class="reference-text"><cite id="CITEREFKulmerMowlaee2015" class="citation conference cs1">Kulmer, Josef; Mowlaee, Pejman (April 2015). "Harmonic phase estimation in single-channel speech enhancement using von Mises distribution and prior SNR". <i>Acoustics, Speech and Signal Processing (ICASSP), 2015 IEEE International Conference on</i>. IEEE. pp.&nbsp;<span class="nowrap">5063–</span>5067.</cite></span>
</li>
<li id="cite_note-16"><span class="mw-cite-backlink"><b><a href="#cite_ref-16">^</a></b></span> <span class="reference-text"><cite id="CITEREFKulmerMowlaee2015" class="citation journal cs1">Kulmer, Josef; Mowlaee, Pejman (May 2015). "Phase Estimation in Single Channel Speech Enhancement Using Phase Decomposition". <i>IEEE Signal Processing Letters</i>. <b>22</b> (5): <span class="nowrap">598–</span>602. <a href="Bibcode_(identifier)" class="mw-redirect" title="Bibcode (identifier)">Bibcode</a>:<a rel="nofollow" class="external text" href="https://ui.adsabs.harvard.edu/abs/2015ISPL...22..598K">2015ISPL...22..598K</a>. <a href="Doi_(identifier)" class="mw-redirect" title="Doi (identifier)">doi</a>:<a rel="nofollow" class="external text" href="https://doi.org/10.1109%2FLSP.2014.2365040">10.1109/LSP.2014.2365040</a>. <a href="ISSN_(identifier)" class="mw-redirect" title="ISSN (identifier)">ISSN</a>&nbsp;<a rel="nofollow" class="external text" href="https://search.worldcat.org/issn/1070-9908">1070-9908</a>. <a href="S2CID_(identifier)" class="mw-redirect" title="S2CID (identifier)">S2CID</a>&nbsp;<a rel="nofollow" class="external text" href="https://api.semanticscholar.org/CorpusID:15503015">15503015</a>.</cite></span>
</li>
<li id="cite_note-Advances-17"><span class="mw-cite-backlink">^ <a href="#cite_ref-Advances_17-0"><sup><i><b>a</b></i></sup></a> <a href="#cite_ref-Advances_17-1"><sup><i><b>b</b></i></sup></a></span> <span class="reference-text"><cite id="CITEREFMowlaeeSaeidiStylianou2016" class="citation journal cs1">Mowlaee, Pejman; Saeidi, Rahim; Stylianou, Yannis (July 2016). <span class="id-lock-subscription" title="Paid subscription required"><a rel="nofollow" class="external text" href="http://linkinghub.elsevier.com/retrieve/pii/S0167639316300784">"Advances in phase-aware signal processing in speech communication"</a></span>. <i>Speech Communication</i>. <b>81</b>: <span class="nowrap">1–</span>29. <a href="Doi_(identifier)" class="mw-redirect" title="Doi (identifier)">doi</a>:<a rel="nofollow" class="external text" href="https://doi.org/10.1016%2Fj.specom.2016.04.002">10.1016/j.specom.2016.04.002</a>. <a href="ISSN_(identifier)" class="mw-redirect" title="ISSN (identifier)">ISSN</a>&nbsp;<a rel="nofollow" class="external text" href="https://search.worldcat.org/issn/0167-6393">0167-6393</a>. <a href="S2CID_(identifier)" class="mw-redirect" title="S2CID (identifier)">S2CID</a>&nbsp;<a rel="nofollow" class="external text" href="https://api.semanticscholar.org/CorpusID:17409161">17409161</a><span class="reference-accessdate">. Retrieved <span class="nowrap">2017-12-03</span></span>.</cite></span>
</li>
</ol></div></div>
<div class="navbox-styles"><style data-mw-deduplicate="TemplateStyles:r1129693374">
/* start https://en.wikipedia.org/ */


.mw-parser-output .hlist dl,.mw-parser-output .hlist ol,.mw-parser-output .hlist ul{margin:0;padding:0}.mw-parser-output .hlist dd,.mw-parser-output .hlist dt,.mw-parser-output .hlist li{margin:0;display:inline}.mw-parser-output .hlist.inline,.mw-parser-output .hlist.inline dl,.mw-parser-output .hlist.inline ol,.mw-parser-output .hlist.inline ul,.mw-parser-output .hlist dl dl,.mw-parser-output .hlist dl ol,.mw-parser-output .hlist dl ul,.mw-parser-output .hlist ol dl,.mw-parser-output .hlist ol ol,.mw-parser-output .hlist ol ul,.mw-parser-output .hlist ul dl,.mw-parser-output .hlist ul ol,.mw-parser-output .hlist ul ul{display:inline}.mw-parser-output .hlist .mw-empty-li{display:none}.mw-parser-output .hlist dt::after{content:": "}.mw-parser-output .hlist dd::after,.mw-parser-output .hlist li::after{content:" · ";font-weight:bold}.mw-parser-output .hlist dd:last-child::after,.mw-parser-output .hlist dt:last-child::after,.mw-parser-output .hlist li:last-child::after{content:none}.mw-parser-output .hlist dd dd:first-child::before,.mw-parser-output .hlist dd dt:first-child::before,.mw-parser-output .hlist dd li:first-child::before,.mw-parser-output .hlist dt dd:first-child::before,.mw-parser-output .hlist dt dt:first-child::before,.mw-parser-output .hlist dt li:first-child::before,.mw-parser-output .hlist li dd:first-child::before,.mw-parser-output .hlist li dt:first-child::before,.mw-parser-output .hlist li li:first-child::before{content:" (";font-weight:normal}.mw-parser-output .hlist dd dd:last-child::after,.mw-parser-output .hlist dd dt:last-child::after,.mw-parser-output .hlist dd li:last-child::after,.mw-parser-output .hlist dt dd:last-child::after,.mw-parser-output .hlist dt dt:last-child::after,.mw-parser-output .hlist dt li:last-child::after,.mw-parser-output .hlist li dd:last-child::after,.mw-parser-output .hlist li dt:last-child::after,.mw-parser-output .hlist li li:last-child::after{content:")";font-weight:normal}.mw-parser-output .hlist ol{counter-reset:listitem}.mw-parser-output .hlist ol>li{counter-increment:listitem}.mw-parser-output .hlist ol>li::before{content:" "counter(listitem)"\a0 "}.mw-parser-output .hlist dd ol>li:first-child::before,.mw-parser-output .hlist dt ol>li:first-child::before,.mw-parser-output .hlist li ol>li:first-child::before{content:" ("counter(listitem)"\a0 "}


/* end https://en.wikipedia.org/ */
</style><style data-mw-deduplicate="TemplateStyles:r1236075235">
/* start https://en.wikipedia.org/ */


.mw-parser-output .navbox{box-sizing:border-box;border:1px solid #a2a9b1;width:100%;clear:both;font-size:88%;text-align:center;padding:1px;margin:1em auto 0}.mw-parser-output .navbox .navbox{margin-top:0}.mw-parser-output .navbox+.navbox,.mw-parser-output .navbox+.navbox-styles+.navbox{margin-top:-1px}.mw-parser-output .navbox-inner,.mw-parser-output .navbox-subgroup{width:100%}.mw-parser-output .navbox-group,.mw-parser-output .navbox-title,.mw-parser-output .navbox-abovebelow{padding:0.25em 1em;line-height:1.5em;text-align:center}.mw-parser-output .navbox-group{white-space:nowrap;text-align:right}.mw-parser-output .navbox,.mw-parser-output .navbox-subgroup{background-color:#fdfdfd}.mw-parser-output .navbox-list{line-height:1.5em;border-color:#fdfdfd}.mw-parser-output .navbox-list-with-group{text-align:left;border-left-width:2px;border-left-style:solid}.mw-parser-output tr+tr>.navbox-abovebelow,.mw-parser-output tr+tr>.navbox-group,.mw-parser-output tr+tr>.navbox-image,.mw-parser-output tr+tr>.navbox-list{border-top:2px solid #fdfdfd}.mw-parser-output .navbox-title{background-color:#ccf}.mw-parser-output .navbox-abovebelow,.mw-parser-output .navbox-group,.mw-parser-output .navbox-subgroup .navbox-title{background-color:#ddf}.mw-parser-output .navbox-subgroup .navbox-group,.mw-parser-output .navbox-subgroup .navbox-abovebelow{background-color:#e6e6ff}.mw-parser-output .navbox-even{background-color:#f7f7f7}.mw-parser-output .navbox-odd{background-color:transparent}.mw-parser-output .navbox .hlist td dl,.mw-parser-output .navbox .hlist td ol,.mw-parser-output .navbox .hlist td ul,.mw-parser-output .navbox td.hlist dl,.mw-parser-output .navbox td.hlist ol,.mw-parser-output .navbox td.hlist ul{padding:0.125em 0}.mw-parser-output .navbox .navbar{display:block;font-size:100%}.mw-parser-output .navbox-title .navbar{float:left;text-align:left;margin-right:0.5em}body.skin--responsive .mw-parser-output .navbox-image img{max-width:none!important}@media print{body.ns-0 .mw-parser-output .navbox{display:none!important}}


/* end https://en.wikipedia.org/ */
</style></div><div role="navigation" class="navbox" aria-labelledby="Computer_audition21" style="padding:3px"><table class="nowraplinks mw-collapsible autocollapse navbox-inner" style="border-spacing:0;background:transparent;color:inherit"><tbody><tr><th scope="col" class="navbox-title" colspan="2"><style data-mw-deduplicate="TemplateStyles:r1239400231">
/* start https://en.wikipedia.org/ */


.mw-parser-output .navbar{display:inline;font-size:88%;font-weight:normal}.mw-parser-output .navbar-collapse{float:left;text-align:left}.mw-parser-output .navbar-boxtext{word-spacing:0}.mw-parser-output .navbar ul{display:inline-block;white-space:nowrap;line-height:inherit}.mw-parser-output .navbar-brackets::before{margin-right:-0.125em;content:"[ "}.mw-parser-output .navbar-brackets::after{margin-left:-0.125em;content:" ]"}.mw-parser-output .navbar li{word-spacing:-0.125em}.mw-parser-output .navbar a>span,.mw-parser-output .navbar a>abbr{text-decoration:inherit}.mw-parser-output .navbar-mini abbr{font-variant:small-caps;border-bottom:none;text-decoration:none;cursor:inherit}.mw-parser-output .navbar-ct-full{font-size:114%;margin:0 7em}.mw-parser-output .navbar-ct-mini{font-size:114%;margin:0 4em}html.skin-theme-clientpref-night .mw-parser-output .navbar li a abbr{color:var(--color-base)!important}@media(prefers-color-scheme:dark){html.skin-theme-clientpref-os .mw-parser-output .navbar li a abbr{color:var(--color-base)!important}}@media print{.mw-parser-output .navbar{display:none!important}}


/* end https://en.wikipedia.org/ */
</style><div id="Computer_audition21" style="font-size:114%;margin:0 4em"><a href="Computer_audition" title="Computer audition">Computer audition</a></div></th></tr><tr><td colspan="2" class="navbox-list navbox-odd hlist" style="width:100%;padding:0"><div style="padding:0 0.25em">
<ul><li><a href="Acoustic_fingerprint" title="Acoustic fingerprint">Acoustic fingerprint</a></li>
<li><a href="Audio_mining" title="Audio mining">Audio mining</a></li>
<li><a href="Computational_auditory_scene_analysis" title="Computational auditory scene analysis">Computational auditory scene analysis</a></li>
<li><a href="Music_information_retrieval" title="Music information retrieval">Music information retrieval</a></li>
<li><a href="Semantic_audio" title="Semantic audio">Semantic audio</a></li>
<li>
<ul><li><a href="Speech_analytics" title="Speech analytics">Speech analytics</a></li>
<li><a href="Speaker_recognition" title="Speaker recognition">Speaker recognition</a></li>
<li><a href="Speech_recognition" title="Speech recognition">Speech recognition</a></li></ul></li>
<li><a href="Sound_recognition" title="Sound recognition">Sound recognition</a></li>
<li><a href="3D_sound_localization" title="3D sound localization">3D sound localization</a></li>
<li><a href="3D_sound_reconstruction" title="3D sound reconstruction">3D sound reconstruction</a></li></ul>
</div></td></tr></tbody></table></div>
<div class="navbox-styles"></div><div role="navigation" class="navbox authority-control" aria-labelledby="Authority_control_databases_frameless&amp;#124;text-top&amp;#124;10px&amp;#124;alt=Edit_this_at_Wikidata&amp;#124;link=https&amp;#58;//www.wikidata.org/wiki/Q3358061#identifiers&amp;#124;class=noprint&amp;#124;Edit_this_at_Wikidata727" style="padding:3px"><table class="nowraplinks hlist mw-collapsible autocollapse navbox-inner" style="border-spacing:0;background:transparent;color:inherit"><tbody><tr><th scope="col" class="navbox-title" colspan="2"><div id="Authority_control_databases_frameless&amp;#124;text-top&amp;#124;10px&amp;#124;alt=Edit_this_at_Wikidata&amp;#124;link=https&amp;#58;//www.wikidata.org/wiki/Q3358061#identifiers&amp;#124;class=noprint&amp;#124;Edit_this_at_Wikidata727" style="font-size:114%;margin:0 4em">Authority control databases </div></th></tr><tr><th scope="row" class="navbox-group" style="width:1%">National</th><td class="navbox-list-with-group navbox-list navbox-odd" style="width:100%;padding:0"><div style="padding:0 0.25em"><ul><li><span class="uid"><a rel="nofollow" class="external text" href="https://id.loc.gov/authorities/sh85126450">United States</a></span></li><li><span class="uid"><a rel="nofollow" class="external text" href="https://id.ndl.go.jp/auth/ndlna/00576334">Japan</a></span></li><li><span class="uid"><a rel="nofollow" class="external text" href="https://www.nli.org.il/en/authorities/987007565841605171">Israel</a></span></li></ul></div></td></tr><tr><th scope="row" class="navbox-group" style="width:1%">Other</th><td class="navbox-list-with-group navbox-list navbox-even" style="width:100%;padding:0"><div style="padding:0 0.25em"><ul><li><span class="uid"><a rel="nofollow" class="external text" href="https://lux.collections.yale.edu/view/concept/33bf8e68-5f76-487e-aee4-8fa2c3fcde9a">Yale LUX</a></span></li></ul></div></td></tr></tbody></table></div></div><!--htdig_noindex--><div><div class="zim-footer">
This article is issued from <a class="external text" title="Last edited on 2025-07-18" href="https://en.wikipedia.org/wiki/?title=Speech_processing&amp;oldid=1301219187">Wikipedia</a>. The text is available under <a class="external text" href="https://creativecommons.org/licenses/by-sa/4.0/deed.en">Creative Commons Attribution-Share Alike 4.0</a> unless otherwise noted. Additional terms may apply for the media files.
</div>
</div><!--/htdig_noindex--></div>
</div>
</main>
</div>
</div>
</div>

</body></html>